Papers with tri-modal architecture

1 papers
Investigating Audio, Video, and Text Fusion Methods for End-to-End Automatic Personality Prediction (P18-2)

Copied to clipboard

Challenge: Using stacked Convolutional Neural Networks, we can predict personality traits from video clips with different channels for audio, text, and video data.
Approach: They propose a tri-modal architecture to predict Big Five personality trait scores from video clips with different channels for audio, text, and video data.
Outcome: The proposed model outperforms the best individual modality with 9.4% accuracy over the best channel.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations